Back

Cell Systems

Elsevier BV

Preprints posted in the last 30 days, ranked by how well they match Cell Systems's content profile, based on 201 papers previously published here. The average preprint has a 0.20% match score for this journal, so anything above that is already an above-average fit.

1
Reliable single-cell perturbations explain and improve model performance

Wang, X.; Kuipers, J.; Hugi, F.; Platt, R. J.; Beerenwinkel, N.

2026-08-12 bioinformatics 10.64898/2026.08.11.744177 medRxiv
Top 0.1%
51.0%
Show abstract

Predicting single-cell transcriptional responses to perturbations is central to building the virtual cell, yet recent benchmarks show that simple baseline methods often outperform complex models, and model comparisons depend on the evaluation metric. Most studies assume that preprocessed RNA sequencing data are reliable ground truth for both training and evaluation. Here, we test this assumption by measuring the reliability of perturbations and their alignment with shared perturbation responses, classifying each perturbation as specific, shared, or unreliable. Among 7,170 perturbations from 29 datasets, 65% are unreliable, 11% shared, and 24% specific. Applying these quality labels to published benchmarks shows that model comparisons depend on perturbation quality. Training with reliable perturbations alone matches or outperforms full-data performance while using 55% of all training perturbations. Our framework also enables prospective experimental design: for most perturbations, a 28-cell pilot experiment accurately predicts how many cells a full screen needs to be reliable.

2
Time-resolved operator archetypes characterize dynamical sensitivity during cell-state transitions

Redd, D. M.; Green, S. G.; Terooatea, T. W.

2026-08-23 bioinformatics 10.64898/2026.08.21.745996 medRxiv
Top 0.1%
32.8%
Show abstract

During development, cells traverse gene expression states where their local dynamical sensitivity changes sharply, yet existing computational methods provide limited access to when and where this sensitivity peaks along a trajectory. Here we introduce scJDO (single-cell Jacobian Differential Operators), a framework that characterizes how local dynamical sensitivity evolves during cell fate transitions, together with an explicit account of what that representation can and cannot recover from snapshot data. scJDO treats time-indexed Jacobians as explicit analytical objects, projecting the temporal sequence of operators into a shared subspace and decomposing it into recurrent operator archetypes with interpretable temporal activation profiles. Unlike methods that derive Jacobians from splicing-kinetic vector fields, scJDO learns a neural drift field directly from cell-state geometry via diffusion score matching, enabling Jacobian analysis on trajectory-resolved scRNA-seq datasets regardless of splicing-data availability. Applied to a dense time-course of induced pluripotent stem cell (iPSC) reprogramming, scJDO resolves a quantitative operator-level signature that distinguishes diverted from productive fate: the productive trajectory executes a sequential handoff from an early MEF-exit operator regime to a late pluripotency-associated regime, whereas the diverted trajectory maintains the early regime and instead activates a distinct stress-associated archetype. We validate scJDO across four settings: synthetic benchmarks with analytically known ground truth, branching hematopoiesis, dense real time-course reprogramming, and a perturbational setting using Schrodinger bridges in K562 CRISPRi Perturb-seq. We compare against the two most widely used single-cell Jacobian methods on a dataset where all three are runnable, finding that scJDO shares significantly more gene-level and directional operator structure with Dynamo than expected by chance while providing operator-level analysis on datasets without splicing kinetics. We further characterize the boundary of the representation directly. At a fate-decision saddle, eight mathematically distinct readouts of the same learned drift field are consistent with a single explanation: a drift field fit to snapshot density reproduces density-dominant separation between committed branches rather than the low-variance transverse instability that defines the decision. Together, scJDO provides an operator-level view of single-cell dynamics and an explicit characterization of its own identifiability boundary.

3
Virtual-cell models compress unseen intervention geometry through a target-specific generalization bottleneck

Huang, Y.; Wang, H.; Wilson, P. C.

2026-08-24 bioinformatics 10.64898/2026.08.21.746243 medRxiv
Top 0.1%
31.3%
Show abstract

Predictive models of cellular perturbation are often judged by how closely they reconstruct molecular states after unseen interventions. We show that high state-level similarity can coexist with loss of the relationships that distinguish perturbations, a failure we term Intervention Geometry Compression (IGC). Across established models and perturbation settings, unseen interventions show weakened global and local geometry, reduced between-intervention variance and spectral collapse. The failure is not primarily explained by response-space capacity. Instead, diagnostic projections localize much of the missing geometry to a small number of residual response directions learned from seen interventions; these directions outperform complexity-matched random subspaces and replicate in an independent Jiang perturbation resource. Polarity captures part, but not all, of this continuous orientation signal. Time-resolved analyses further show that correct trajectory entry markedly improves downstream propagation, while a held target's own early empirical response rapidly reveals endpoint orientation. Finally, same-target empirical anchoring transfers intervention identity across contexts far more effectively than increasing exposure to other interventions. These results identify intervention-coordinate assignment as an information bottleneck in virtual-cell generalization and support a design principle: empirically anchor intervention identity, then use models to generalize anchored effects across cellular contexts.

4
Double Machine Learning with Multi-Gene Shared Backgroundfor Causal Inference in Single-Cell Data: Grouping Deviation Follows a Random Walk and the Accuracy-Compute Trade-Off

Ye, W.; Jiang, X.; Shen, F.

2026-08-10 systems biology 10.64898/2026.08.08.743701 medRxiv
Top 0.1%
27.8%
Show abstract

In high-throughput single-cell transcriptomics (p {approx} 20,000 genes), performing double machine learning (DML) causal inference on q {approx} 5,000 target genes requires nuisance function fits that grow linearly with the number of targets (Kf cross-fitting folds, Kf = 5 or 10), far exceeding feasible computational budgets, especially with deep learning. We propose a Randomized Partition Strategy (RPS): randomly divide target genes into groups, share one background compression per group, reducing deep learning model training to q/m runs (m = group size) --- a factor of m savings. The cost of grouping is accuracy loss --- we prove that the cumulative deviation of the estimator follows a one-dimensional drift-free symmetric random walk, with diffusion variance growing linearly with group size and mean squared displacement equaling the mean squared error, so accuracy loss is predictable: m = 1 is always optimal, accuracy cost is monotonically increasing, and a small accuracy sacrifice yields m-fold compute savings. On GSE189050 SLE single-cell data (Memory B cells, n = 2120), both PCA and DL methods converge to the same conclusion, confirming the random walk mechanism is method-independent; an unexpected finding is that DL diffusion growth is only 16%, far slower than PCAs 7.4 times. This work provides a quantifiable theoretical foundation for compute strategy selection in single-cell high-dimensional causal inference.

5
Learning protein function through autonomous experimental interaction

Brooks, C.; Notin, P.; Romero, P. A.

2026-08-20 bioengineering 10.64898/2026.08.14.744985 medRxiv
Top 0.1%
27.2%
Show abstract

Biological AI learns primarily from existing observations, but many questions cannot be answered from available data alone. Here we show that AI can instead acquire knowledge by acting directly on biological systems and learning from the consequences. We developed a closed-loop framework in which autonomous agents design protein variants, construct and characterize them in a robotic laboratory, learn from the resulting experimental feedback, and decide what experiments to perform next. We then allowed the system to operate continuously and without human intervention for approximately one month, during which multiple agents independently explored protein sequence space while learning from shared experimental experience. Applied to glycoside hydrolases, the agents discovered enzymes with substantially altered substrate specificity toward non-native sugars and progressively learned the structure of the underlying sequence-function landscape. The resulting experimental experience also revealed determinants of substrate specificity and protein expression that were not specified as learning objectives. These results demonstrate that AI can autonomously interact with biology over extended periods to acquire knowledge through experience, establishing a framework for biological discovery driven by continuous experimental interaction.

6
A permutation-free family-wise error rate for the moderated top-gene scan under gene correlation

Dwyer, W. J.

2026-08-21 bioinformatics 10.64898/2026.08.17.745282 medRxiv
Top 0.1%
26.9%
Show abstract

A differential-expression scan reports the genes with the largest moderated t-statistics, so controlling the family-wise error rate means controlling the null distribution of the maximum statistic over genes. Under gene correlation this is widely believed to require permutation, because correlation changes the effective multiplicity and corrupts the empirical-Bayes variance prior behind the moderated t-statistic. We decompose that liberality by an error-budget ablation and show that, within the simulated model class, it reduces principally to an inflation of the empirical-Bayes prior degrees of freedom: substituting the true prior returns the family-wise error to the independent-gene small-sample baseline, so dependence imposes no separate barrier once the prior is correct. Correlation deflates the cross-gene spread of the log sample variances; because the prior degrees of freedom decreases in that spread, the prior is over-estimated and the moderated maximum turns liberal. Dividing the observed spread by one minus the mean squared gene correlation, estimated by a tuning-free spectral U-statistic with an unbiased trace target, reverses the mechanism and holds the family-wise error near the baseline at retained power without permutation. An observable instability index flags when severe co-expression should defer to permutation.

7
MSGPCA: Multi-Slice Graph PCA for replicate-aware Spatial Omics analysis

Chakraborty, A.; Neelon, B.; Lawson, A.; Angel, P.; Chung, D.; Seal, S.

2026-08-11 bioinformatics 10.64898/2026.08.05.743056 medRxiv
Top 0.1%
26.8%
Show abstract

As spatial transcriptomics (ST) and spatial proteomics (SP) technologies mature, experimental designs are increasingly moving beyond single-slice analyses toward multi-slice studies involving one or more donors and experimental conditions. Although these designs enable the identification of reproducible spatial signals, they also introduce substantial biological heterogeneity, particularly when integrating non-serial slices or anatomically distinct regions. If not modeled carefully, such variation can blur slice-specific tissue structure, mask conserved molecular patterns, and limit the discovery of biologically relevant latent structure. Although dimension reduction is essential for representing high-dimensional molecular data in a lower-dimensional space, existing multi-slice methods typically enforce a globally shared representation that inadequately accommodates slice-level heterogeneity. To address this limitation, we propose Multi-Slice Graph Principal Component Analysis (MSGPCA), which decomposes molecular variation into shared spatial factors conserved across slices and slice-specific factors that capture local tissue microarchitecture. In downstream analyses, MSGPCA-derived representations recover spatial tissue structure, denoise molecular profiles, and reveal biologically interpretable metafeatures associated with shared and slice-specific biology. In a mass spectrometry imaging dataset comprising nonserial slices of ductal carcinoma in situ (DCIS) and invasive breast cancer (IBC), the shared factors captured broad biological differences across tissue regions, whereas the slice-specific factors revealed intratumoral spatial variation within the IBC microenvironment. In human dorsolateral prefrontal cortex ST data, MSGPCA recovered laminar cortical architecture across adjacent slices, closely aligning with expert pathologist annotations. Together, these findings demonstrate that MSGPCA resolves shared tissue architecture while preserving local microenvironmental variation in complex multi-slice spatial omics datasets.

8
Microenvironment-informed inference of transcriptional progression geometry

Kobara, S.; Rahman, S. A.; Ribeiro, S. P.; Coopersmith, C. M.; Kamaleswaran, R.

2026-08-24 bioinformatics 10.64898/2026.08.21.746284 medRxiv
Top 0.1%
26.3%
Show abstract

We present BIOCURRENT, a causal inference framework that reconstructs donor-specific pseudotime geometry in transcriptomic data. By modeling gene expression as a function of baseline characteristics, microenvironmental context, and latent pseudotime, BIOCURRENT enables comparison of compressed or expanded progression intervals across transcriptional state transitions. We introduce $\Delta\Delta T$, a geometry-based estimator that quantifies differences in pseudotime intervals across conditions, enabling evaluation of changes in pseudotime intervals under hypothetical modulation of microenvironmental programs. Applications to thymic T-cell developmental lineages and to COVID-19 immune dysregulation reveal condition- and donor-specific distortions of progression intervals. Counterfactual simulation links microenvironmental context to changes in specific intracellular state transition intervals. By localizing deviations in pseudotime geometry, BIOCURRENT identifies whether shifts in transcriptomic programs emerge early or later along transcriptomic coordinates and reveals upstream programs associated with these distortions. Such localization supports transcriptional stage-aware mechanistic hypotheses and suggests candidate intervention checkpoints in complex biological systems.

9
Assay concordance sets exact ceilings on what one biological score can predict

Liu, Z.

2026-08-26 bioinformatics 10.64898/2026.08.24.746774 medRxiv
Top 0.1%
21.8%
Show abstract

Computational models of biology are ranked by averaging one prediction against many experimental realizations of a phenotype that are treated as interchangeable. We show this imposes an exact, model-free ceiling fixed by how much those realizations agree with each other, and that the ceiling depends on the evaluation metric through a single support-function identity. Measuring assay concordance across four public registries, 2,822 MaveDB score sets, 217 ProteinGym assays, two drug screens and 1,150 CRISPR cell lines, we find that two assays of one target agree at 0.56-0.68, and that 541 domains measured twice with different proteases fix assay reliability at 0.897, so 70-90% of every ceiling is irreducible biology rather than noise. Published predictors realize 63% of the achievable on the correlation benchmarks report and 18% on the top-1% selection their users perform. We provide the estimator, the ceilings, and the measurements the field has not made.

10
A generalized growth law for translation- and transcription-targeting antibiotics captures drug interactions

Gadjisade, N.; Mori, M.; Bollenbach, T.

2026-08-24 systems biology 10.64898/2026.08.21.746212 medRxiv
Top 0.1%
19.2%
Show abstract

Bacterial growth laws quantitatively connect intracellular resource allocation to growth rate, enabling accurate predictions of physiology and antibiotic responses. Yet these laws have been rigorously tested for only a handful of perturbations. Here, we show that the growth law linking ribosome levels to growth rate under translation-inhibiting antibiotics is not universal, but rather depends on the antibiotic's mechanism of action. Quantitative proteomics across finely resolved one- and two-dimensional antibiotic gradients showed that inhibitors of translocation elongation or peptide bond formation elicit the canonical rise in ribosome levels, consistent with the growth law. By contrast, antibiotics disrupting translation initiation or fidelity produced distinct responses without ribosome upregulation. The transcription inhibitor rifampicin even reduced ribosome abundance. Combining antibiotics with divergent ribosome responses revealed a generalized growth law, in which the individual responses to perturbations superimpose. Embedding this law in a mathematical model explains distinct drug interaction patterns observed between rifampicin and different translation inhibitors. A low-dimensional structure pervades the entire proteome, enabling prediction of responses to drug pairs based on single-drug measurements. Together, these findings broaden the scope of bacterial growth laws and provide new principles for predicting responses to antibiotic combinations.

11
Symbolic regression enables coarse-grained model discovery of intracellular signalling dynamics

de Pomereu, T.; Fröhlich, F.

2026-08-21 systems biology 10.64898/2026.08.20.745973 medRxiv
Top 0.1%
19.1%
Show abstract

Cells respond to their environment through protein networks often dysregulated in cancer, making dynamical modelling crucial. Limitations in experimental data and computational resources motivate coarse-graining methods to build low-dimensional descriptions. Yet classical approaches to coarse-grained modelling rely on strong assumptions, leaving it unclear when partial experimental observations support reduced descriptions of system dynamics. Here we show that symbolic regression (SR) provides a data-driven way to test whether, and how compactly, the dynamics of a signalling system coarse-grain over the measured variables, and, when they do, infers mechanistically interpretable models. In synthetic enzyme systems, SR recovers Michaelis-Menten kinetics for the two-step mechanism and under three-step extensions. As data quality is degraded, SR simplifies toward effective kinetic laws while preserving correct theoretical limits. Applied to published time-resolved ERK phosphorylation data, SR identifies compact phospho-ERK rate laws in selected cancer-relevant gene overexpression contexts, yielding interpretable kinetic effects. A sparse neural ODE baseline requires few inputs where SR succeeds, but on average more where it fails, indicating that, where a reduced model is learnable at all, SR failure is associated with more complex dynamics that a simple mathematical model cannot describe. Together, these findings establish symbolic regression as a way to test when a compact coarse-grained description is warranted, generating hypotheses where one holds and motivating potential new measurements where it does not.

12
Conditional source attribution at plankton bloom onset: Identifiability and sharp bounds under environmental forcing

Caputi, L.

2026-08-18 ecology 10.64898/2026.08.13.744670 medRxiv
Top 0.1%
18.8%
Show abstract

Can observations distinguish a bloom supplied from within a study volume from one supplied across its boundary? We develop a theoretical framework for that question at plankton bloom onset, conditional on a predeclared, observed or calibrated onset event and a declared set of environmental paths, biological responses, and model forms. The estimand follows source-event labels through forcing-dependent survival and genotype-specific growth. Its central certificate asks whether the local onset fraction is invariant over every source history that produces the same time-expanded observation record. For polyhedral history fibers, a Charnes-Cooper transformation computes both sharp dynamic-data endpoints as linear programs. When each source instead has a fixed normalized onset signature, the certificate reduces to a row-space test; uncertain signatures require a joint lifted program. For a finite compatible scenario ensemble, admissible fractions are the union across scenarios, and a point is justified only when every nonempty scenario gives the same singleton. A synthetic two-genotype witness gives the same observed total but local fractions of 2/3 and 1/3 under reversed forcing-response gains. The observer, mixture, and optimization ingredients are established; the contribution is their target-specific synthesis around source at onset. The framework is diagnostic rather than predictive. It specifies what a study must measure--local sources, boundary inflow, forcing, response, timing, and carrier signatures on one declared window--and returns an interval when missing components have justified bounds, including [0, 1] when they remain unconstrained.

13
A confound-diagnostic toolkit for in silico perturbation with single-cell foundation models

Qiu, R.; Zhao, M. M.

2026-08-07 bioinformatics 10.64898/2026.08.04.732812 medRxiv
Top 0.1%
18.7%
Show abstract

Deleting a gene token from a cells input sequence offers a convenient native strategy for in silico perturbation, but the resulting embedding delta may not represent a biological knockout response. Apparent effects can instead reflect gene identity, universal responsiveness, limited tokenization coverage, library-size contamination, or circular state scoring. Here, we present a confound-diagnostic framework combining held-out increment testing, responsiveness adjustment, coverage gating, library-size diagnostics, and de-circularized state-shift analysis, together with a numerically matched reimplementation of frozen Geneformers perturbation engine. Across Frangieh and Replogle datasets and linear and nonlinear readouts, the native embedding delta provided no reproducible held-out improvement beyond gene identity. Signal-injection calibration showed that the test detected injected residual signal, whereas native increments remained below its detection floor. Matched controls traced apparent positives to raw-count library-size structure, broad responsiveness, and self-referential scoring, while coverage constrained perturbation applicability and estimate stability without establishing biological specificity. This model-adaptable framework helps determine when foundation-model perturbation readouts warrant biological interpretation. MotivationFoundation-model in silico perturbation could predict perturbation effects when matched experimental data are unavailable. However, in zero-shot settings, embedding-derived responses may reflect gene identity, universal responsiveness, tokenization limits, library-size artifacts, or circular state scoring rather than biological knockout effects. We therefore developed a reusable confound-diagnostic framework that applies matched controls to test whether native perturbation readouts contain information beyond these confounds and warrant biological interpretation.

14
RADF: Reference-Anchored Dynamic Flow for Spatial Perturbation Profile Completion

Cai, H.; Wang, H.; Chen, J.; Xue, Z.; Sheng, X.; Zhang, T.

2026-08-24 bioinformatics 10.64898/2026.08.20.745474 medRxiv
Top 0.1%
18.5%
Show abstract

Spatial perturbation profiling is becoming an important tool in functional genomics because it reveals how genetic interventions reshape transcription within intact tissue contexts. However, destructive readout and limited screening capacity leave many perturbation-by-location response profiles unmeasured, motivating the task of spatial perturbation profile completion. The task is to infer the held-out response population at query locations from reported profiles of the same perturbation. Existing methods either generate responses de novo or reuse these profiles without spatial adaptation. These strategies make it difficult to preserve empirical population structure while modeling location-specific variation. Our key insight is that the reported population already defines an empirical response distribution for the target perturbation. To exploit this empirical support, we propose Reference-Anchored Dynamic Flow (RADF), which employs a Sinkhorn-balanced decoder to construct a population-valued anchor in which every reference profile has equal total contribution. Additionally, a bounded dynamic relational flow is used to recompute spatial relations from the evolving expression state and query geometry. Across diverse spatial contexts, RADF reduces macro E-distance by 70.6% compared with an existing state-of-the-art spatial method, highlighting the advantage of combining a reference-supported population anchor with bounded, location-dependent refinement. Code will be made publicly available upon acceptance.

15
Regulatory stochasticity drives opposing phenotypic outcomes in cell-fate decision networks

Hari, K.; Gupta, A.; Shivakumar, L. M.; Kulkarni, P.; Salgia, R.; Jolly, M. K.; Levine, H.

2026-08-26 systems biology 10.64898/2026.08.24.746658 medRxiv
Top 0.1%
18.4%
Show abstract

Gene regulatory network models treat interaction parameters as fixed, although regulatory efficacy fluctuates. We asked how temporal fluctuations in interaction strength reshape phenotype occupancy in cell-fate decision GRN motifs. Across large parameter ensembles, anchored fluctuations largely preserved deterministic occupancies. Additive fluctuations increased occupancy of all-high co-expression states, particularly where high expression was accessible. In contrast, multiplicative fluctuations biased inhibitory interactions toward stronger repression and favored single-high states in a topology-dependent manner. Deterministic controls sampled from noise-induced parameter distributions did not fully reproduce these effects. A Boolean-limit analysis revealed an intrinsic upward bias: loss of repression increased expression regardless of regulator state, whereas stronger repression acted only when the regulator was present. Analyses of epithelial-mesenchymal plasticity and gonadal-fate networks showed increased occupancy of hybrid team-expression states under additive fluctuations. Thus, regulatory noise can reshape the developmental landscape in opposing directions, pushing cell-fate systems toward either progenitor-like or terminally differentiated states.

16
A mechanism-annotated benchmark reveals limited fidelity to drug-response signatures in single-cell perturbation models

Li, L.; Duan, S.; Zha, X.; Ye, F.; Zhang, Y.; Zhang, X.; Cao, Y.; Liu, C.

2026-08-24 bioinformatics 10.64898/2026.08.19.745729 medRxiv
Top 0.1%
18.2%
Show abstract

Single-cell drug perturbation models are increasingly used to predict how compounds remodel cellular states, but they are still largely assessed by expression reconstruction. Whether high expression similarity reflects preservation of drug-response signatures remains unclear. Here we present scDrugPerturb-Bench, a mechanism-annotated benchmark that links matched control and drug-treated single-cell RNA-sequencing profiles to literature-curated directional key-gene evidence. The resource covers 181 datasets, 423 annotated response cases, 717 unique key genes and 2.5 million cells. We introduce the Mechanism Fidelity Score (MFS) to evaluate key-gene direction, effect-size recovery, gene-set coherence, mechanism specificity and pathway-level response polarity. Across 12 perturbation-prediction models, 3 baselines and 10 data splits, expression-similarity metrics were weakly aligned with MFS and selected different model configurations. Mechanism-aware selection improved early drug retrieval in a transcriptome-based drug design evaluation, indicating that MFS provides practical information beyond benchmark reporting. Systematic benchmarking revealed limited fidelity to drug-response signatures across cell-line and source-integrated settings. Frozen single-cell foundation model embeddings produced local, metric-dependent gains rather than universal improvements, and source context substantially reshaped model assessment. Hard-negative tests further showed that plausible perturbation responses can arise from non-specific transcriptional shortcuts. These results show that expression reconstruction is an insufficient proxy for preserving drug-response signatures and establish scDrugPerturb-Bench as a benchmark for mechanism-aware evaluation of single-cell drug perturbation models.

17
Single-cell foundation models benefit from cross-modal training: adding proteomics data beats parameter scaling

Burq, M.; Stepec, D.; Kim, C.; Cimermancic, P.

2026-08-19 bioinformatics 10.64898/2026.08.14.744845 medRxiv
Top 0.1%
18.1%
Show abstract

Leading cellular foundation models have been trained on hundreds of millions of single-cell transcriptomes, with progress increasingly driven by larger datasets and model scaling. Here, we asked whether adding a proteomics modality can improve gene-level and cell-level representations beyond scaling RNA-only models. We introduce cross-modal continued pretraining, fine-tuning a published single-cell model (Tahoe-x1) on a large corpus of proteomic profiles. Training a 70M-parameter Tahoe-x1 model for a single epoch on 48843 proteomic samples from 440 diverse mass-spectrometry studies matched or exceeded 1B- and 3B-parameter RNA-only models across most of the original Tahoe-x1 evaluation benchmarks. This shows that with the right training recipe, heterogeneous proteomics data can improve the learned representations of single-cell RNAseq samples, demonstrating strong out-of-distribution generalization. Cross-modal pretraining also improves transfer to a held-out protein perturbation benchmark, where scaling the RNA-only model does not provide comparable benefits. These results demonstrate that careful targeted curation of proteomics data can provide larger benefits than increasing the model size alone and suggest that multimodal pretraining is a promising path toward more informative biological foundation models.

18
Benchmark Averages Hide the Failures That Matter: Quantizing ESM-2 for Protein Variant-Effect Prediction

Shao, Q.

2026-08-18 bioinformatics 10.64898/2026.08.10.744024 medRxiv
Top 0.1%
18.0%
Show abstract

We benchmark six numerical precision configurations for ESM-2 protein language models across throughput, memory footprint and predictive accuracy, on two workloads with sharply different characteristics: bulk embedding extraction and deep mutational scanning (DMS) variant-effect scoring. Accuracy is evaluated on the complete ProteinGym substitution benchmark -- 201 assays, 2.41M variants -- at three model scales spanning 650M to 15B parameters, with a paired bootstrap clustered on protein. Three findings follow, and each contradicts a common practice. First, benchmark averages conceal the failure that decides deployability: no configuration shifts mean correlation by more than 0.007 at any scale, yet INT8 dynamic quantization -- indistinguishable from fp32 on that mean at 3B (p = 0.34) -- takes a single assay from{rho} = 0.591 to 0.223. Selection must be made on worst-case, not mean, behaviour. Second, fidelity measured against fp32 bounds risk but cannot rank quality: over 3015 assay/configuration pairs it predicts the magnitude of ground-truth change (r = 0.56-0.81) but not its direction, and the INT4 effect differs significantly between 650M and 3B (+0.0101, p = 0.0007) with no monotone trend to extrapolate. Third, quantizing a large model is dominated by using a small one: of eighteen scale/configuration combinations only three are Pareto-optimal over accuracy, memory and speed, and all three are 650M. The one catastrophic failure we observe is a defect of default symmetric activation scaling, not of W8A8 itself: asymmetric activation quantization, a one-line change needing no calibration, removes every damaged assay. We also give a label-free screen for at-risk targets, and report four measurement artifacts encountered during this study, three of which inverted the result they were meant to measure.

19
A Generative Virtual Tissue Model Enables Computational Design of Therapeutic Perturbation Strategies

Lu, Y.; Zhang, W.; Chen, Y.-J.; Yin, J.; Chen, L.; Fleisher, K.; Gornet, J.; Liu, R.; Wang, Z. J.; Poon, Y.; You, Y.; Thomson, M.

2026-08-20 bioinformatics 10.64898/2026.08.12.743536 medRxiv
Top 0.1%
15.3%
Show abstract

Computational design has transformed many fields of engineering, where simulators can explore millions of candidate design configurations before experimental development and testing. Therapeutic design in biomedicine has resisted computational design approaches because disease progression and therapeutic response emerge from interactions among many cell types within human tissue, governed by biochemical parameters that are largely unknown and potentially unknowable. Here, we introduce the Cell Interaction Foundation Model (CIFM), a virtual tissue model that forward-simulates the transcriptional dynamics of cells in human tissue under arbitrary therapeutic conditions based upon a spatial transcriptomic seed. CIFM is a geometric graph neural network trained by self-supervised masked-transcriptome prediction on millions of cellular microenvironments spanning human tissue types and disease states; generative, auto-regressive, monte-carlo play-out, then, simulates transcriptional dynamics under combinatorial perturbations from a spatial transcriptomic seed. We validate CIFM by showing accuracy gains in gene expression prediction and imputation, disease classification, recapitulation of perturbation responses in prostate cancer models, and recovery of T cell-tumor signaling measured in cell-cell sequencing experiments. Beyond such conventional tasks, CIFM enables target identification and therapeutic design through generative tissue simulation play-outs. Analyzing over 106 single and combinatorial perturbations, CIFM designs immunotherapy strategies for cancer and autoimmune disease that exploit combinatorial manipulation of signaling pathways to induce or suppress immune activation. Broadly, CIFM shows how generative artificial intelligence methods can be applied to model emergent behavior in highly interacting biological systems, yielding new approaches to fundamental understanding of tissue behavior as well as large-scale therapeutic design.

20
Move BeTween modAlities (MBTA) employs flow matching to predict single cell data modalities

Xu, B.; Zhang, Y.; Michor, F.

2026-08-09 bioinformatics 10.64898/2026.08.05.743110 medRxiv
Top 0.1%
15.2%
Show abstract

Integrating diverse molecular modalities to obtain a comprehensive view of cellular identity remains a major challenge in single-cell biology. A fundamental but underappreciated obstacle is structural mismatch -- the phenomenon in which the neighborhood structure of a cell differs depending on which molecular modality is used to define it. Existing approaches typically embed modalities into a shared latent space, which actively erases the structural differences between modalities that make multimodal measurements scientifically valuable. Here we introduce Move BeTween modAlities (MBTA), the first framework explicitly designed to address structural mismatch. Rather than forcing modalities into a shared representation, MBTA maintains modality-specific latent spaces and connects them via flow matching, preserving the structural integrity of each modality while enabling accurate cross-modal translation. Across extensive benchmarks on multi-modal single-cell datasets, MBTA consistently outperformed existing methods, with the largest gains observed in datasets with pronounced structural mismatch. Applied to joint genomic and transcriptomic profiles of breast cancer patients, MBTA identified transcriptomic lineage relationships corroborated by genomic variation and outperformed state-of-the-art transcriptomics-based copy number inference methods. Extending this framework to mouse embryonic development, we reconstructed temporal trajectories jointly defined by gene expression and seven complementary epigenetic modalities. MBTA can connect any number of molecular readouts without erasing their individual character, serving as the computational foundation for assembling multi-layered portraits of cells.